Papers with text-based approaches
Dialogue Act-based Breakdown Detection in Negotiation Dialogues (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent studies have succeeded in modeling a negotiating agent in natural language that can control both text generation and reasoning in goal-oriented dialogue systems. |
| Approach: | They propose a human-human negotiation dialogue dataset that features increased complexities in terms of the number of possible solutions and a utility function. |
| Outcome: | The proposed method performs comparable to text-based approaches in existing corpora and better results in the proposed dataset. |
Multimodal Conversation Modelling for Topic Derailment Detection (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on analysing textual dialogues that derailed into toxic content ignores visual information, such as images and videos. |
| Approach: | They propose a new multimodal conversational architecture that utilises visual and conversational contexts to classify comments for derailment. |
| Outcome: | The proposed approach outperforms existing methods and is more robust to textual noise. |
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)
Copied to clipboard
| Challenge: | despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation . |
| Approach: | They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture . |
| Outcome: | The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs. |
Incorporating Object-Level Visual Context for Multimodal Fine-Grained Entity Typing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Experimental results show that fine-grained entity typing is superior to text-based methods. |
| Approach: | They propose a task called fine-grained entity typing to classify entities . they propose combining textual and visual contexts to capture fine-granular semantic information . |
| Outcome: | The proposed approach achieves superior classification performance compared to previous text-based approaches. |
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval (2025.acl-long)
Copied to clipboard
| Challenge: | Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. |
| Approach: | They propose a visual-textual embedding framework that integrates textual and visual features for robust document representation. |
| Outcome: | The proposed visual-textual embedding framework surpasses existing methods while preserving semantic fidelity. |